Papers with rule-based system
Scalable and Robust Self-Learning for Skill Routing in Large-Scale Conversational AI Systems (2022.naacl-industry)
Copied to clipboard
| Challenge: | Existing methods to enable skill routing do not scale in terms of the number of skills and skill on-boarding. |
| Approach: | They propose a model-based approach to enable natural conversation by allowing frequent policy updates . they propose an annotation-based system, rule-based model, and bandit-based learning . |
| Outcome: | The proposed method is scalable and cost-effective, the authors show . they show that it can improve the user experience without abrupt policy changes . |
A Simple Unsupervised Approach for Coreference Resolution using Rule-based Weak Supervision (2022.starsem-1)
Copied to clipboard
| Challenge: | state-of-the-art coreference models rely on labeled data, but an end-to-end model is needed to solve this problem. |
| Approach: | They propose an approach that leverages an end-to-end neural model in settings where labeled data is unavailable. |
| Outcome: | The proposed approach outperforms the previous best unsupervised model and outperformed the rule-based model on English OntoNotes corpus. |
Team SVMrank: Leveraging Feature-rich Support Vector Machines for Ranking Explanations to Elementary Science Questions (D19-53)
Copied to clipboard
| Challenge: | TextGraphs 2019 Shared Task on Multi-Hop Inference for Explanation Regeneration tackles explanation generation for elementary science questions. |
| Approach: | They propose a hybrid pipelined machine learning model and rule-based system to address MIER-19 . they use a featurerich learning-to-rank machine learning and a rule-driven system to rerank the LTR model predictions. |
| Outcome: | The proposed model was ranked fourth in the evaluation, close to the second and third ranked teams, achieving 39.4% MAP. |
Exploring Interpretability in Event Extraction: Multitask Learning of a Neural Event Classifier and an Explanation Decoder (2020.acl-srw)
Copied to clipboard
| Challenge: | EE is a key requirement for machine learning in many domains, e.g., legal, medical, finance. |
| Approach: | They propose an interpretable approach for event extraction that jointly trains a classifier and a rule decoder for event processing. |
| Outcome: | The proposed approach can be used for semi-supervised learning and its performance improves when trained on automatically-labeled data generated by a rule-based system. |
Emotion Cause Extraction on Social Media without Human Annotation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have focused on extracting emotion causes from news articles, but lack of fine-grained annotations has limited the ECE task. |
| Approach: | They propose a new ECE framework that extracts emotion causes from social media data without relying on human annotations. |
| Outcome: | The proposed framework achieves high extraction performance and generalizability without relying on human annotations. |
Summarizing Patients’ Problems from Hospital Progress Notes Using Pre-trained Sequence-to-Sequence Models (2022.coling-1)
Copied to clipboard
| Challenge: | Problem list summarization requires a model to understand, abstract, and generate clinical documentation. |
| Approach: | They propose a task that summarises patients' main problems from daily progress notes using input from the provider's progress notes during hospitalization. |
| Outcome: | The proposed model outperforms two state-of-the-art seq2seq transformer architectures in summarizing patients' main problems from daily progress notes in the medical information mart for Intensive Care (MIMIC)-III. |
Enhancing Sequence-to-Sequence Neural Lemmatization with External Resources (2021.eacl-main)
Copied to clipboard
| Challenge: | a hybrid approach to lemmatization enhances the seq2seq neural model with additional lemmas extracted from an external lexicon or a rule-based system. |
| Approach: | They propose a hybrid approach that enhances a seq2seq neural model with additional lemmas extracted from an external lexicon or a rule-based system. |
| Outcome: | The proposed model achieves statistically significant improvements on 23 UD languages, compared to baseline models not utilizing additional lemma information. |
HarfoSokhan: A Comprehensive Parallel Dataset for Transitions between Persian Colloquial and Formal Variations (2026.eacl-long)
Copied to clipboard
Hamid Jahad Sarvestani, Vida Ramezanian, Saee Saadat, Neda Taghizadeh Serajeh, Maryam Sadat Razavi Taheri, Shohreh Kasaei, MohammadAmin Fazli, Ehsaneddin Asgari
| Challenge: | A wide array of NLP/NLU models have been developed for the Persian language but performance drops when applied to the colloquial form of Persian. |
| Approach: | They propose to use a large-scale colloquial to formal Persian parallel dataset to train a GPT2 model that exhibited remarkable proficiency in colloqual to informal text style transfer. |
| Outcome: | The proposed dataset outperforms OpenAI’s GPT-3.5-turbo model and a leading rule-based system in colloquial to formal Persian conversion. |
Harnessing Pre-Trained Neural Networks with Rules for Formality Style Transfer (D19-1)
Copied to clipboard
| Challenge: | Existing studies normalize informal sentences with rules, but they introduce noise if we use them in a naive way. |
| Approach: | They propose to harness rules into a state-of-the-art neural network that is typically pretrained on massive corpora. |
| Outcome: | The proposed method can be used to generate a state-of-the-art on a small dataset. |
Intrinsic Task-based Evaluation for Referring Expression Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Referring Expression Generation (REG) models generate referring expressions that refer to referents at different points in a discourse. |
| Approach: | They propose to use a purely ratings-based human evaluation to evaluate REG models by completing two meta-level tasks. |
| Outcome: | The proposed evaluation makes the models more reliable and discriminable, and improves the quality of the REs. |
GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns (2025.findings-acl)
Copied to clipboard
| Challenge: | Gender rewriting is an NLP task that uses gendered forms to mitigate gender biases. |
| Approach: | They propose a French gender-neutral rewriting system using collective nouns, which are gender-fixed in French. |
| Outcome: | The proposed system detects gendered forms and replaces them with neutral or opposite forms. |
The Use of Text Alignment in Semi-Automatic Error Analysis: Use Case in the Development of the Corpus of the Latvian Language Learners (L18-1)
Copied to clipboard
| Challenge: | Using error annotation methods, the corpus of the Latvian language learners can be adapted for other languages with relatively free word order. |
| Approach: | They propose a method for creating error annotated corpora using text correction, automated morphological analysis, automated text alignment and error annotation. |
| Outcome: | The proposed method has been approbated in the development of the corpus of the Latvian language learners. |
Replace and Report: NLP Assisted Radiology Report Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Clinical practice frequently uses medical imaging for diagnosis and treatment. |
| Approach: | They propose a template-based approach to generate radiology reports from radiographs . they use multilabel image classifiers to generate tags, pathological descriptions from tags . |
| Outcome: | The proposed method improves on the most popular radiology report datasets. |
Enhancing Accessible Communication: from European Portuguese to Portuguese Sign Language (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing systems for translating European Portuguese into LGP glosses rely on hand-crafted rules . current systems rely only on toy examples, disregarding non-manual movements . |
| Approach: | They propose a corpora-driven rule-based machine translation system between European Portuguese and LGP glosses and two neural machine translation models. |
| Outcome: | The proposed system improves on existing translation systems and annotates a gold collection of the results. |
PictoEduca: Building a Dataset for Spanish Text-to-Pictogram Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | PictoEduca is the first large-scale Spanish text-to-pictogram dataset for augmentative and alternative communication. |
| Approach: | They present PictoEduca, a large-scale Spanish text-to-pictogram dataset for augmentative and alternative communication. |
| Outcome: | The proposed dataset combines automatic annotation with targeted expert correction, supporting scalable and high-quality corpus construction. |